Papers with computer vision

96 papers
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track (2025.emnlp-industry)

Copied to clipboard

Challenge: EMNLP 2025 Industry Track highlights key insights, novel research trends and challenges encountered in practical language technology applications.
Approach: Kai Chen will present the technical advances behind the open-source Intern-series large models . he will highlight how models acquire expert-level skills in specialized domains .
Outcome: This talk will highlight the technical advances behind the open-source Intern-series models . it will highlight how models acquire expert-level skills in specialized domains while retaining broad generalization ability.
Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (2022.naacl-main)

Copied to clipboard

Challenge: Silver es high quality training datasets for d AI training data labelilng services academic institutions engaged in ce R&D and application research to processing (NLP), voice recognition hesis (TTS), and computer vision (CV).
Approach: Silver es provides high quality training datasets for d AI training data labelilng services academic institutions engaged in ce R&D and application research to processing (NLP), voice recognition hesis (TTS), and computer vision (CV).
Outcome: Silver es high quality training datasets for d AI training data labelilng services academic institutions engaged in ce R&D and application research to processing (NLP), voice recognition hesis (TTS), and computer vision (CV).
ISA: An Intelligent Shopping Assistant (2020.aacl-demo)

Copied to clipboard

Challenge: In-store users only need to take a picture or scan the barcode of the product of interest, and then the user can talk to the assistant about the product.
Approach: They present a mobile-based intelligent shopping assistant that is designed to improve shopping experience in physical stores.
Outcome: The proposed system can improve shopping experience in physical stores by leveraging advanced techniques in computer vision, speech processing, and natural language processing.
Explanation in the Era of Large Language Models (2024.naacl-tutorials)

Copied to clipboard

Challenge: Explanation has long been a part of communication, where humans use language to elucidate each other and transmit information about mechanisms of events.
Approach: They review the opportunities and challenges of explanations in the era of large language models and examine how they can be used to generate explanations.
Outcome: The proposed methods are based on the models of large language models (LLMs) and their opaque nature.
Transfer Learning in Natural Language Processing (N19-5)

Copied to clipboard

Challenge: supervised machine learning is based on learning in isolation, a single predictive model for a task using a dataset.
Approach: They present an overview of modern transfer learning methods in natural language processing . they review examples and case studies on how models can be integrated and adapted .
Outcome: The proposed methods improve upon the state-of-the-art on a wide range of NLP tasks.
Pushing the Limits of Radiology with Joint Modeling of Visual and Textual Information (P18-3)

Copied to clipboard

Challenge: Recent research has focused on the intersection of computer vision and natural language processing, but its adaption to the medical domain is not fully explored.
Approach: They aim to develop machine learning models that can reason jointly on medical images and clinical text for advanced search, retrieval, annotation and description of medical images.
Outcome: The proposed models can reason jointly on medical images and clinical text for advanced search, retrieval, annotation and description of medical images.
Recognizing Multimodal Entailment (2021.acl-tutorials)

Copied to clipboard

Challenge: This tutorial introduces the multimodal entailment task for detecting semantic alignments . the task requires fine-grained understanding of visual and linguistic semantics questions .
Approach: This tutorial introduces the multimodal entailment task to machine learning . it introduces a dataset for recognizing multimodal alignments .
Outcome: This tutorial introduces the multimodal entailment task . it can be useful for detecting semantic alignments when a single modality alone is not enough .
Batch-Softmax Contrastive Loss for Pairwise Sentence Scoring Tasks (2022.naacl-main)

Copied to clipboard

Challenge: Recent advances in machine learning have led to the use of contrastive loss for representation learning.
Approach: They propose to use batch-softmax contrastive loss to train pairwise sentence embeddings . they propose to take a batch-softermax contrastitive loss and train it with different loss functions .
Outcome: The proposed model improves on a number of datasets and pairwise sentence scoring tasks.
Generating Image Captions in Arabic using Root-Word Based Recurrent Neural Networks and Deep Neural Networks (N18-4)

Copied to clipboard

Challenge: Existing studies on image caption generation in English focus on Western languages, ignoring Semitic and Middle-Eastern languages like Arabic, Hebrew, Urdu and Persian.
Approach: They propose to leverage the critical dependency of Arabic to generate Arabic captions using root-word based Recurrent Neural Network and Deep Neural networks.
Outcome: The proposed model outperforms English-Arabic translated captions on a dataset from newspapers in the Middle East.
Advancing the Robustness of Large Language Models through Self-Denoised Smoothing (2024.naacl-short)

Copied to clipboard

Challenge: Existing adversarial attacks can cause LLMs to make wrong predictions on downstream tasks or generate harmful content misaligned with human values.
Approach: They propose to use randomized smoothing to add noise to the input and then make predictions based on these denoised versions.
Outcome: The proposed method surpasses existing methods in both empirical and certified robustness in defending against adversarial perturbations for both downstream tasks and human alignments (i.e., jailbreak attacks).
Kandinsky: An Improved Text-to-Image Synthesis with Image Prior and Latent Diffusion (2023.emnlp-demo)

Copied to clipboard

Challenge: Experimental evaluations demonstrate FID score of 8.03 on the COCO-30K dataset, marking our model as the top open source performer in terms of measurable image generation quality.
Approach: They propose a latent diffusion-based model that combines image prior and latent diffusive techniques to create a text-to-image architecture.
Outcome: The proposed model achieves the highest FID score among open-source models . it is compared with the state-of-the-art models on the COCO-30K dataset .
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels (2024.eacl-long)

Copied to clipboard

Challenge: Obtaining high-quality labeled data that accurately represents complexity of real-world scenarios can be expensive, time-consuming, or even impractical.
Approach: They propose to use Fréchet Inception Distance to measure distance between judged items and retrieved results.
Outcome: The proposed method improves on a MS MARCO dataset and TREC Deep Learning Tracks query sets.
Semi-supervised Adversarial Text Generation based on Seq2Seq models (2022.emnlp-industry)

Copied to clipboard

Challenge: In contrast, adversarial training has been used in computer vision to improve models’ robustness due to the discrete nature of text.
Approach: They propose a way to generate adversarial samples by using pseudo-labeled in-domain text data to train a seq2seq model for adversarials and combine it with paraphrase detection.
Outcome: The proposed model generates realistic and relevant adversarial samples compared to other state-of-the-art models and recovers up to 70% of errors.
Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies (2025.findings-naacl)

Copied to clipboard

Challenge: Audio descriptions (ADs) are acoustic commentaries designed to assist blind and visually impaired individuals in accessing digital media content.
Approach: They examine how state-of-the-art NLP and CV technologies can be applied to generate ADs . they identify essential research directions for the future .
Outcome: The proposed technologies can be applied to generate audio descriptions (ADs) the process is time-consuming and costly, and requires significant human effort . the authors identify key research directions for the future .
Foundation Model for Biomedical Graphs: Integrating Knowledge Graphs and Protein Structures to Large Language Models (2024.acl-srw)

Copied to clipboard

Challenge: Transformer model has been a de-facto standard in natural language processing, but it is limited to images, text, and/or sequence data.
Approach: They propose to use a multimodal large language model architecture to handle biomedical graphs such as protein structure and chemical molecules to improve its performance.
Outcome: The proposed architecture can handle multiple data types for biomedical graphs such as protein structure and chemical molecules.
Universal Language Model Fine-tuning for Text Classification (P18-1)

Copied to clipboard

Challenge: Existing approaches to computer vision require task-specific modifications and training from scratch.
Approach: They propose a method that can be applied to any task in NLP and propose to open-source it.
Outcome: The proposed method outperforms the state-of-the-art on six text classification tasks, reducing error by 18-24% on majority of datasets.
Parameter-free and Accessible Prompt Learning to Enhance Adversarial Robustness for Pre-trained Vision-Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large pre-trained Vision-Language Models (VLMs) have revolutionized downstream vision-language tasks including classification, object detection, and segmentation.
Approach: They propose to search for text prompts at the word level rather than optimizing continuous textual embeddings to boost adversarial robustness.
Outcome: Experiments show that the proposed method outperforms hand-engineered prompts with average gains of +4.9% and +5.8%.
ConceptBert: Concept-Aware Representation for Visual Question Answering (2020.findings-emnlp)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is a challenging task that has received increasing attention from both the computer vision and the natural language processing communities.
Approach: They propose an algorithm which learns a joint Concept-Vision-Language embedding for questions which require common sense knowledge from external structured content.
Outcome: The proposed model is based on the Outer Knowledge-VQA and VQA datasets.
PaperMage: A Unified Toolkit for Processing, Representing, and Manipulating Visually-Rich Scientific Documents (2023.emnlp-demo)

Copied to clipboard

Challenge: Existing tools for working with scientific documents are limited and documents are often in difficult-to-use PDF formats.
Approach: They propose an open-source Python toolkit for analyzing and processing visually-rich scientific documents.
Outcome: PaperMage provides turn-key recipes for common scientific document processing use-cases.
DropMix: A Textual Data Augmentation Combining Dropout with Mixup (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods to overcome overfitting in text learning do not consider dimensionality . dimensionalization is important for deep neural networks to overcome the problem .
Approach: They propose a saliency map-based approach to overcome overfitting in text learning . they propose augmentation regularization methods such as Dropout and Mixup to improve regularization .
Outcome: Empirical results show that the proposed approach overcomes overfitting in text learning . dropout and mixup methods are effective in enhancing regularization .
Multimodal Pretraining Unmasked: A Meta-Analysis and a Unified Framework of Vision-and-Language BERTs (2021.tacl-1)

Copied to clipboard

Challenge: Large-scale pretraining and task-specific fine-tuning are now the standard methodology for many tasks in computer vision and natural language processing.
Approach: They propose to combine two types of vision and language BERTs to create a theoretical framework that can be unified under different theoretical frameworks.
Outcome: The proposed models can be classified into single-stream or dual-stream encoders and are unified under a single theoretical framework.
LM-CPPF: Paraphrasing-Guided Data Augmentation for Contrastive Prompt-Based Few-Shot Fine-Tuning (2023.acl-short)

Copied to clipboard

Challenge: Recent advances in pre-trained language models have been limited when fine-tuned on small datasets.
Approach: They propose to add contrastive learning to prompt-based fine-tuning to improve model performance.
Outcome: The proposed approach outperforms other methods on multiple text classification benchmarks.
Teacher Intervention: Improving Convergence of Quantization Aware Training for Ultra-Low Precision Transformers (2023.eacl-main)

Copied to clipboard

Challenge: Quantization-aware training (QAT) is a promising method to lower the implementation cost and energy consumption.
Approach: They propose a method for fast converging QAT of pre-trained Transformers using a layer-wise signal propagation method with the intact signal from the teacher.
Outcome: The proposed method achieves superior accuracy with significantly lower fine-tuning iterations on well-known Transformers of natural language processing as well as computer vision compared to the state-of-the-art methods.
Two Methods for Domain Adaptation of Bilingual Tasks: Delightfully Simple and Broadly Applicable (P18-1)

Copied to clipboard

Challenge: Previously, domain adaptation approaches to bilingual tasks were proposed . we show that simple adaptation process involving only unlabeled text is highly effective .
Approach: They propose a method for domain adaptation of bilingual word embeddings using unlabeled data . they then tailor a semi-supervised classification method from computer vision to these tasks .
Outcome: The proposed method improves on two bilingual tasks using unlabeled data.
The Hidden Attention of Mamba Models (2025.acl-long)

Copied to clipboard

Challenge: Recent studies have shown that Mamba models can be used for multiple domains, including NLP, long-range sequence processing, and computer vision.
Approach: They add a third view and show that Mamba models can be viewed as attention-driven models.
Outcome: The proposed model can be viewed as attention-driven and empirically compare it to the attention-based models of transformers.
Effective Few-Shot Classification with Transfer Learning (2020.coling-main)

Copied to clipboard

Challenge: Recent work on few-shot learning addresses the problem of learning based on a small amount of training data.
Approach: They adapt the Amazon Review Sentiment Classification (ARSC) text dataset for few-shot learning . they train a single binary classifier to learn all few- shot classes jointly .
Outcome: The proposed approach outperforms most published results on the ARSC text dataset . the results suggest that the classes in the AR SC few-shot task are very similar to each other .
Self-Governing Neural Networks for On-Device Short Text Classification (D18-1)

Copied to clipboard

Challenge: Existing deep neural networks have a tiny memory footprint and low computational capacity compared to high performance computing systems such as CPUs, GPUs and TPUs on the cloud.
Approach: They propose on-device self-governing neural networks which learn compact projection vectors with local sensitive hashing.
Outcome: The proposed models perform better on dialog act classification tasks while maintaining high accuracy.
Label Anchored Contrastive Learning for Language Understanding (2022.naacl-main)

Copied to clipboard

Challenge: a novel approach to contrastive learning for language understanding is not fully explored . contrastive training has been widely applied to self-supervised representation learning .
Approach: They propose a label anchored contrastive learning approach for language understanding using a class label.
Outcome: The proposed approach improves on GLUE and CLUE benchmarks by 4.1% compared to the state-of-the-art approaches . the proposed approach also improves under the few-shot and data imbalance settings .
Self-Governing Neural Networks for On-Device Short Text Classification (D18-1)

Copied to clipboard

Challenge: Existing deep neural networks have a tiny memory footprint and low computational capacity compared to high performance computing systems such as CPUs, GPUs and TPUs on the cloud.
Approach: They propose on-device self-governing neural networks which learn compact projection vectors with local sensitive hashing.
Outcome: The proposed models perform better on dialog act classification tasks while maintaining high accuracy.
LiViTo: Linguistic and Visual Features Tool for Assisted Analysis of Historic Manuscripts (2020.lrec-1)

Copied to clipboard

Challenge: a mixed methods approach is feasible for the identification of scribes and authors in handwritten documents.
Approach: They propose a mixed methods approach to the identification of scribes and authors in handwritten documents . they use a software tool which combines linguistic insights and computer vision techniques .
Outcome: The proposed tool can be used to identify scribes and authors in handwritten documents.
PDF-to-Text Reanalysis for Linguistic Data Mining (L18-1)

Copied to clipboard

Challenge: In the 1990s, extracting semistructured text from scientific writing in PDF files was largely a computer vision and OCR problem.
Approach: They propose a system for the reanalysis of PDF-extracted text that performs block detection, respacing, and tabular data analysis for linguistic data mining.
Outcome: The proposed system eliminates the extreme verbosity of XML output while leaving important positional information available for downstream processes.
Achilles-Bench: A Challenging Benchmark for Low-Resource Evaluation (2024.findings-acl)

Copied to clipboard

Challenge: Existing low-resource datasets that challenge neural networks cause over-estimated performance, despite promising yet saturated results in high-res settings.
Approach: They propose a benchmark Achilles-Bench to better evaluate the learning ability of neural networks in low-resource settings.
Outcome: The proposed benchmarks show that even pre-trained language models show performance drops on NLP tasks.
Evaluating the Faithfulness of Importance Measures in NLP by Recursively Masking Allegedly Important Tokens and Retraining (2022.findings-emnlp)

Copied to clipboard

Challenge: To explain NLP models, importance measures are often used to inform input tokens are important for making a prediction.
Approach: They propose a faithfulness metric that masks allegedly important tokens and retrains the model.
Outcome: The proposed metric is based on LSTM-attention models and RoBERTa models.
Exploring Data Augmentation in Neural DRS-to-Text Generation (2024.eacl-long)

Copied to clipboard

Challenge: Neural networks are notoriously data-hungry, resulting in ungrammatical texts . data augmentation requires a specific design for a structurally rich input format .
Approach: They propose to selectively augment a training set with new data by adding and varying two specific lexical categories, i.e. proper and common nouns.
Outcome: The proposed approach selectively augments a training set with new data by adding and varying two specific lexical categories, i.e. proper and common nouns.
Language is All a Graph Needs (2024.findings-eacl)

Copied to clipboard

Challenge: Existing work on integrating graph problems into generative language modeling framework remains limited.
Approach: They propose an LLM with instructions based on natural language to perform graph tasks.
Outcome: The proposed model surpasses all GNN baselines on ogbn-arxiv, Cora and PubMed datasets and sheds light on generative LLMs as new foundation model for graph machine learning.
Few-shot Learning for Slot Tagging with Attentive Relational Network (2021.eacl-main)

Copied to clipboard

Challenge: Recent studies have used metric-based learning in computer vision but not slot tagging.
Approach: They propose a metric-based learning architecture that extends relation networks by leveraging pretrained contextual embeddings such as ELMO and BERT and by using attention mechanism.
Outcome: The proposed method outperforms state-of-the-art methods on SNIPS data on a slot tagging task with a large amount of hand-labeled data.
bert2BERT: Towards Reusable Pretrained Language Models (2022.acl-long)

Copied to clipboard

Challenge: Pre-training large language models can be expensive and wasteful.
Approach: They propose a method which can transfer the knowledge of an existing smaller pre-trained model to a large model through parameter initialization and a two-stage learning method to further accelerate the pre-training.
Outcome: The proposed method can transfer the knowledge of an existing smaller pre-trained model to a large model through parameter initialization and significantly improve the pre-training efficiency of the large model.
Question Modifiers in Visual Question Answering (2022.lrec-1)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is a multi-disciplinary task that requires integration of several key disciplines.
Approach: They develop a model that adds modifiers to questions based on object properties and spatial relationships using Amazon Mechanical Turk data.
Outcome: The proposed model can improve when questions are modified to include more details.
Locally Aggregated Feature Attribution on Natural Language Model Understanding (2022.naacl-main)

Copied to clipboard

Challenge: a growing popularity of deep-learning models makes model understanding more important . feature attribution methods have shown promising results in computer vision but are not trivial .
Approach: They propose a gradient-based feature attribution method that smooths gradients by aggregating similar reference texts derived from language model embeddings.
Outcome: The proposed method outperforms existing methods on public datasets and key words detection tasks.
Automated Extraction of Prosodic Structure from Unannotated Sign Language Video (2024.lrec-main)

Copied to clipboard

Challenge: a new method for analyzing prosody in sign languages uses the velocity profile of the hands . the velocity profiles of hand movements can be used to analyse prosodic structure .
Approach: They propose a method for extracting velocity information from unlabeled video of sign language using CoTracker.
Outcome: The proposed method can extract prosodic information from unlabeled video clips.
End-to-End Construction of NLP Knowledge Graph (2021.findings-acl)

Copied to clipboard

Challenge: a new schema for NLP knowledge about tasks, datasets and metrics is proposed.
Approach: They propose a new schema that represents knowledge about tasks, datasets and metrics in the NLP domain.
Outcome: The proposed framework can be automatically built into scientific leaderboards . the proposed system achieves reasonable results for all relation types on this small-scale graph .
Improving Generation and Evaluation of Visual Stories via Semantic Consistency (2021.naacl-main)

Copied to clipboard

Challenge: Story visualization is an underexplored task that requires a generative model to generate images . prior work has focused on image generation but there is room for improvement .
Approach: They propose to add a dual learning framework to reinforce semantic alignment between story and generated images and a copy-transform mechanism to model sequentially-consistent story visualization.
Outcome: The proposed models outperform text-to-image synthesis models on the story visualization task . the proposed models also improve visual quality, coherence and relevance .
Talk2Car: Taking Control of Your Self-Driving Car (D19-1)

Copied to clipboard

Challenge: a long-term goal of artificial intelligence is to have an agent execute commands through natural language.
Approach: They propose to use a dataset to compare commands written in natural language for self-driving cars with other datasets.
Outcome: The proposed task is a challenging one and shows promising results, the authors argue . the talk2car dataset compares with similar datasets and shows that the proposed task requires additional research in natural language processing and computer vision.
SumCSE: Summary as a transformation for Contrastive Learning (2024.findings-naacl)

Copied to clipboard

Challenge: Sentence embedding models are typically trained using contrastive learning (CL) using human annotations directly or by repurposing other annotated datasets.
Approach: They propose to use generative language models to generate CL data using annotated data.
Outcome: The proposed method outperforms the previous best unsupervised method by 1.8 points and SimCSE, a strong supervised baseline by 0.3 points on the semantic text similarity (STS) benchmark.
The Impact of Differential Privacy on Group Disparity Mitigation (2024.findings-naacl)

Copied to clipboard

Challenge: a recent study evaluated the impact of differential privacy on fairness across four tasks.
Approach: They evaluate the impact of differential privacy on fairness across four diverse tasks . they train (,)-differentially private models with empirical risk minimization .
Outcome: The proposed model shows that differential privacy increases performance differences between groups . the model also reduces performance differences in the robust setting .
A Survey of Retentive Network (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on the effectiveness of the Retentive Networks have not yet been conducted.
Approach: They propose a retention mechanism that integrates the inductive bias of recurrent neural networks with the parallelizable training advantages of attention-based models.
Outcome: The proposed retention mechanism combines the inductive bias of recurrent neural networks with the parallelizable training advantages of attention-based models.
Visual Prompting in LLMs for Enhancing Emotion Recognition (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for enhancing in-context emotion classification fail to include spatial relationships between different people and facial features within a single face.
Approach: They propose a set-of-vision prompting approach that uses spatial information to mark targets precisely.
Outcome: The proposed approach improves face count and emotion categorization while preserving the enriched image context.
Challenges in Pre-Training Graph Neural Networks for Context-Based Fake News Detection: An Evaluation of Current Strategies and Resource Limitations (2024.lrec-main)

Copied to clipboard

Challenge: Graph Neural Networks (GNNs) are used to train neural networks to detect fake news based on context-based methods.
Approach: They propose to combine the two by applying pre-training of Graph Neural Networks (GNNs) in the domain of context-based fake news detection.
Outcome: The proposed methods show that transfer learning does not lead to significant improvements over training a model from scratch in the domain of context-based fake news detection.
CHICA: A Developmental Corpus of Child-Caregiver’s Face-to-face vs. Video Call Conversations in Middle Childhood (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies of language-in-interaction focus on the two ends of the developmental spectrum, i.e., early childhood and adulthood, leaving a gap in our knowledge about how development unfolds, especially across middle childhood.
Approach: They propose to use CHICA to analyze child-caregiver conversations at home . they use mobile, lightweight eye-tracking and head motion detection to optimize the naturalness of the recordings.
Outcome: The proposed corpus of child-caregiver conversations at home was compared with a previous corpus based on a set of conversations between children aged 7, 9, and 11 years old.
SupCL-Seq: Supervised Contrastive Learning for Downstream Optimized Sequence Representations (2021.findings-emnlp)

Copied to clipboard

Challenge: SupCL-Seq extends contrastive learning from computer vision to sequence classification tasks.
Approach: They propose a supervised alternative to Masked Language Modeling (MLM) that extends contrastive learning to sequence optimization in NLP by altering the dropout mask probability in standard Transformer architectures.
Outcome: The proposed method leads to large gains on the GLUE benchmark, including 6% absolute improvement on CoLA, 5.4% on MRPC, 4.7% on RTE and 2.6% on STS-B.
Eye4Ref: A Multimodal Eye Movement Dataset of Referentially Complex Situations (2020.lrec-1)

Copied to clipboard

Challenge: Eye4Ref is a rich multimodal dataset of eye-movement recordings from referentially complex situated settings.
Approach: They present a rich multimodal dataset of eye-movement recordings from situated settings . they use linguistic labels, saccadic movement parameters and symbolic knowledge representations .
Outcome: The Eye4Ref dataset is an annotated multimodal dataset from three eyetracking studies on reference resolution and disambiguation tasks in situated settings.
Language Repository for Long Video Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Language-based learning models (LLMs) support long context-lengths but their effectiveness in handling long-term information gradually declines with input length.
Approach: They propose a Language Repository (LangRepo) that maintains concise and structured information as an interpretable representation.
Outcome: The proposed framework is evaluated on zero-shot visual question-answering benchmarks.
Handling Extreme Class Imbalance in Technical Logbook Datasets (2021.acl-long)

Copied to clipboard

Challenge: Technical logbooks are a challenging and under-explored text type in automated event identification.
Approach: They propose a feedback strategy that resamples the training data based on its error in the prediction process.
Outcome: The proposed approach provides the best results for four different neural network models trained across a suite of technical logbook datasets from distinct technical domains.
ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence Embedding (2022.coling-1)

Copied to clipboard

Challenge: a new method for learning unsupervised sentence embeddings is proposed . unsup-SimCSE is biased because of the length information encoded into the sentence embeds .
Approach: They propose a new unsupervised sentence embedding method that uses dropout to obtain positive pairs from a pre-trained Transformer encoder.
Outcome: The proposed method outperforms the state-of-the-art unsup-SimCSE on a STS task.
SlimFit: Memory-Efficient Fine-Tuning of Transformer-based Models Using Training Dynamics (2024.naacl-long)

Copied to clipboard

Challenge: SlimFit reduces the memory requirements of transformer-based models by analyzing their training dynamics and freezing less-contributory layers during fine-tuning.
Approach: They propose a tool that analyzes transformer-based models and freezes less-contributory layers during fine-tuning to reduce the overall on-device memory usage.
Outcome: SlimFit reduces the memory requirements of transformer-based models by analyzing their training dynamics and freezing less-contributory layers during fine-tuning.
How to Determine the Most Powerful Pre-trained Language Model without Brute Force Fine-tuning? An Empirical Survey (2023.findings-emnlp)

Copied to clipboard

Challenge: Transferability estimation has been a topic of great interest in computer vision fields . a lack of a comprehensive comparison between these estimation methods is a problem .
Approach: They conduct a thorough survey of existing methods to find the most suitable model . they also outline difficulties of consideration of training details and applicability to text generation .
Outcome: The proposed methods perform well with superiorities in effectiveness and efficiency.
Improving Generalizability in Implicitly Abusive Language Detection with Concept Activation Vectors (2022.acl-long)

Copied to clipboard

Challenge: a new study shows that general abusive language classifiers are reliable in detecting explicit abuse but fail to detect more subtle abuses.
Approach: They propose an interpretability technique to quantify the sensitivity of a trained model to new data . they propose a degree of explicitness metric to suggest out-of-domain unlabeled examples .
Outcome: The proposed interpretability technique is useful for predicting the generalizability of the model on new data.
Does Gender Matter? Towards Fairness in Dialogue Systems (2020.coling-main)

Copied to clipboard

Challenge: Recent studies have shown that AI is unfair in many real-world applications such as computer vision and recommendations.
Approach: They propose to use a benchmark dataset to study the fairness of dialogue systems to understand their bias.
Outcome: The proposed methods reduce the bias in dialogue systems significantly.
Beyond Embeddings: The Promise of Visual Table in Visual Reasoning (2024.emnlp-main)

Copied to clipboard

Challenge: Visual representation learning has been a cornerstone in computer vision for decades.
Approach: They propose a visual representation tailored for visual reasoning that provides instance-level world knowledge and detailed attributes that are essential for visual reason.
Outcome: The proposed visual tables outperform existing models on 11 visual reasoning benchmarks.
Universal Domain Adaptation for Robust Handling of Distributional Shifts in NLP (2023.findings-emnlp)

Copied to clipboard

Challenge: Despite advances in computer vision, its application on language input still needs to be explored despite its feasibility.
Approach: They propose a universal domain adaptation (uniDA) benchmark for natural language that offers thorough viewpoints of the model’s generalizability and robustness.
Outcome: The proposed model can handle spoken language in the real world while also detecting unprocessable inputs from the target domain.
How Effective is Task-Agnostic Data Augmentation for Pretrained Transformers? (2020.findings-emnlp)

Copied to clipboard

Challenge: Task-agnostic data augmentations have proven widely effective in computer vision, even on pretrained models.
Approach: They examine the effects of two types of task-agnostic data augmentation on pretrained transformers using 5 classification tasks and 6 datasets.
Outcome: The proposed techniques improve performance on 5 classification tasks, 6 datasets, and 3 variants of modern pretrained transformers.
ETAS: Zero-Shot Transformer Architecture Search via Network Trainability and Expressivity (2024.findings-acl)

Copied to clipboard

Challenge: Existing Transformer Architecture Search methods are limited to computer vision and natural language processing tasks.
Approach: They propose a Transformer Architecture Search proxy that measures trainability and expressivity of Transformer networks separately and integrates it into an effective regularized evolution framework to demonstrate its efficacy.
Outcome: The proposed proxy can achieve higher correlation with the true performance of Transformer networks on computer vision and natural language processing tasks.
Visuo-Linguistic Question Answering (VLQA) Challenge (2020.findings-emnlp)

Copied to clipboard

Challenge: Understanding images and text together is an important aspect of cognition and building advanced AI systems.
Approach: They propose to derive joint inference about a given image-text modality and compile a question-answering corpus using an image and a reading passage.
Outcome: The proposed method has better baseline performance but is still far behind human performance.
DET: A Dual-Encoding Transformer for Relational Graph Embedding (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to graph representation only consider the local neighbors, sacrificing the Transformer’s ability to attend to elements at any distance.
Approach: They propose a dual-encoding Transformer architecture that uses a structural encoder and a semantic encoder to seek for semantically relevant nodes.
Outcome: The proposed architecture achieves superior performance compared to state-of-the-art attention-based methods on complex relational graphs like KGs and citation networks.
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation (2025.findings-naacl)

Copied to clipboard

Challenge: Text-to-image generation models exhibit a strong bias toward English-speaking cultures, ignoring or misrepresenting the unique characteristics of other language groups, countries, and nationalities.
Approach: They propose a RusCode benchmark to evaluate the quality of text-to-image generation containing elements of the Russian cultural code.
Outcome: The proposed model is based on 1250 text prompts in Russian and their translations into English.
SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies focus on the text modality or are limited to specific tasks.
Approach: They propose a framework to teach Large Vision-Language Models to selectively utilize retrieved information and improve their robustness against irrelevant or misleading references.
Outcome: The proposed framework improves LVLMs’ ability to utilize retrieved multimodal references and their robustness against irrelevant or misleading information.
On Vision Features in Multimodal Machine Translation (2022.acl-long)

Copied to clipboard

Challenge: Recent work on multimodal machine translation (MMT) has focused on the way of incorporating vision features into translation but little attention is given to the quality of vision models.
Approach: They develop a selective attention model to study the patch-level contribution of an image in multimodal machine translation.
Outcome: The proposed model is able to learn translation from the visual modality on probing tasks and is compared with existing models.
DiffusionDialog: A Diffusion Model for Diverse Dialog Generation with Latent Space (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have tried to introduce discrete or Gaussian-based latent variables to address the one-to-many problem, but the diversity is limited.
Approach: They propose a diffusion model to enhance the diversity of dialogue generation by using continuous latent variables instead of discrete ones.
Outcome: The proposed model greatly enhances diversity of dialog response while keeping the coherence.
PerturbScore: Connecting Discrete and Continuous Perturbations in NLP (2023.findings-emnlp)

Copied to clipboard

Challenge: Natural language processing (NLP) applications are growing rapidly due to discrete nature of texts.
Approach: They propose to connect discrete perturbations with continuous perturbations to help understand discrete ones in NLP models.
Outcome: The proposed method surpasses methods used in discrete perturbation measuring and can be generalized to different datasets, perturbation methods.
Bridge the Gap Between CV and NLP! A Gradient-based Textual Adversarial Attack Framework (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for adversarial samples are poorly applied in computer vision . however, textual adversarials are still vulnerable to small perturbations .
Approach: They propose a framework to extend existing adversarial attack methods to textual adversarials by adding optimized perturbations to embedding layer and amplifying them in forward propagation process.
Outcome: The proposed framework achieves better performance even using proxy gradient information and produces more fluent and grammatical adversarial samples compared to baseline methods.
DiffuseDef: Improved Robustness to Adversarial Attacks via Iterative Denoising (2025.acl-long)

Copied to clipboard

Challenge: Existing adversarial defense methods for natural language processing still pose challenges to adversarials.
Approach: They propose a novel adversarial defense method that incorporates a diffusion layer as a denoiser between the encoder and the classifier.
Outcome: The proposed method improves over existing adversarial defense methods and achieves state-of-the-art performance against black-box and white-box adversarials.
Instance-adaptive training with noise-robust losses against noisy labels (2021.emnlp-main)

Copied to clipboard

Challenge: Several noise-robust losses have been proposed and evaluated on tasks in computer vision, but they use a single dataset-wise hyperparamter to control the strength of noise resistance.
Approach: They propose to change single dataset-wise hyperparameters of noise resistance to be instance-wise.
Outcome: The proposed frameworks increase noise-robustness on noisy and corrupted NLP datasets.
HoLLMwood: Unleashing the Creativity of Large Language Models in Screenwriting via Role Playing (2024.findings-emnlp)

Copied to clipboard

Challenge: Generative AI has demonstrated unprecedented creativity in the field of computer vision, yet such phenomena have not been observed in the realm of literary creation.
Approach: They propose a framework for unleashing the creativity of large language models (LLMs) they assign LLMs to different roles involved in real-world scenario, they write .
Outcome: The proposed framework outperforms baselines in terms of coherence, relevance, interestingness and overall quality on automatically generated screenplays.
Symmetric Dot-Product Attention for Efficient Training of BERT Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Transformer-based models are stretched to enormous sizes, requiring increasingly larger training datasets and unsustainable amount of compute resources.
Approach: They propose an alternative compatibility function for the Transformer-based attention mechanism that exploits an overlap in the learned representation of the traditional scaled dot-product attention mechanism.
Outcome: The proposed model achieves 79.36 on the GLUE benchmark against 78.74 for the traditional implementation and reduces the number of trainable parameters by 6%.
VideoQA-TA: Temporal-Aware Multi-Modal Video Question Answering (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for video question answering align visual or textual features directly with large language models, limiting the deep semantic association between modalities and hindering a comprehensive understanding of interactions within spatial and temporal contexts.
Approach: They propose a temporal-aware framework for multi-modal video question answering that aligns videos and questions at fine-grained levels.
Outcome: The proposed framework improves reasoning ability and accuracy of videoQA by aligning videos and questions at fine-grained levels.
EFTNAS: Searching for Efficient Language Models in First-Order Weight-Reordered Super-Networks (2024.lrec-main)

Copied to clipboard

Challenge: Depending on the size of transformer-based models, they can be restricted from deployment in resource-constrained environments.
Approach: They propose to combine neural architecture search and network pruning techniques to generate and train weight-sharing super-networks that contain efficient transformer-based models.
Outcome: The proposed model achieves high-performing, high-performance subnetworks on the general language understanding evaluation and the Stanford Question Answering Dataset.
Transkimmer: Transformer Learns to Layer-wise Skim (2022.acl-long)

Copied to clipboard

Challenge: Prior work has proposed to augment Transformer model with the capability of skimming tokens to improve its computational efficiency.
Approach: They propose to add a parameterized predictor before each layer that learns to make the skimming decision.
Outcome: The proposed model achieves 10.97x speedup on GLUE benchmark compared with BERT-base baseline with less than 1% accuracy degradation.
DAPE V2: Process Attention Score as Feature Map for Length Extrapolation (2025.acl-long)

Copied to clipboard

Challenge: Extensive experiments demonstrate that treating attention as a feature map and applying convolution as . a processing method significantly enhances Transformer performance.
Approach: They propose to use the convolution operator to mimic the processing methods in computer vision to treat attention as a feature map and apply it to neighboring attention scores across different heads.
Outcome: The proposed model can be adapted to various attention-related models and achieves high performance.
Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions (2022.acl-long)

Copied to clipboard

Challenge: Vision-and-Language Navigation (VLN) is a research topic that is gaining attention in the field of artificial intelligence.
Approach: They propose to build an embodied agent that can communicate with humans in natural language and navigate in real 3D environments.
Outcome: This paper reviews current studies in the emerging field of vision-and-language navigation . it highlights limitations and opportunities for future work .
“That Is a Suspicious Reaction!”: Interpreting Logits Variation to Detect NLP Adversarial Attacks (2022.acl-long)

Copied to clipboard

Challenge: Existing methods to detect adversarial text inputs are limited in performance and are not detectable via spell checkers.
Approach: They propose a model-agnostic detector of adversarial text examples that detects patterns in the logits of the target classifier when perturbing the input text.
Outcome: The proposed detector improves the state-of-the-art performance in recognizing adversarial inputs and exhibits strong generalization capabilities across different NLP models, datasets, and word-level attacks.
Depth Growing for Neural Machine Translation (P19-1)

Copied to clipboard

Challenge: Neural machine translation models with tens and even more than a hundred blocks have shown effectiveness in image recognition.
Approach: They propose a two-stage approach with three specially designed components to construct deeper NMT models.
Outcome: The proposed approach improves on WMT14 EnglishGerman and EnglishFrench translation tasks.
Context-aware Interactive Attention for Multi-modal Sentiment and Emotion Analysis (D19-1)

Copied to clipboard

Challenge: Multi-modal analysis is a field emerging in the fields of natural language processing, computer vision and speech processing . multimodal analysis uses a variety of information from multiple sources to build efficient systems . acoustic and visual information can provide better information for classification decisions .
Approach: They propose a recurrent neural network based approach for multi-modal sentiment and emotion analysis . they employ a context-aware attention module to exploit the correspondence among neighboring utterances .
Outcome: The proposed model learns inter-modal interaction among participating modalities through auto-encoder mechanism . it is compared with existing state-of-the-art models on five standard multi-modal affect analysis datasets .
Evaluating Saliency Explanations in NLP by Crowdsourcing (2024.lrec-main)

Copied to clipboard

Challenge: a crowdsourced method to evaluate saliency methods in NLP is proposed . saliencies are difficult for humans to understand, and can cause psychological harm .
Approach: They propose a method to evaluate saliency methods in NLP by crowdsourcing . they recruited 800 crowd workers and empirically evaluated seven salience methods .
Outcome: The proposed method evaluates saliency methods on two datasets using crowdsourced data . it shows that the results are comparable to existing methods on NLP and CV fields .
Few-shot Temporal Pruning Accelerates Diffusion Models for Text Generation (2024.lrec-main)

Copied to clipboard

Challenge: Existing acceleration methods for text generation ignore the importance of the distribution of sampling steps, resulting in slow sampling rates.
Approach: They propose a technique to accelerate diffusion models for text generation without additional training by using a Bayesian optimization approach.
Outcome: The proposed technique achieves 400x acceleration even with minimal sampling steps after down to less than 1 minute of optimization yielding a competitive performance even with minimum sampling steps.
A Generative Pre-Trained Language Model for Channel Prediction in Wireless Communications Systems (2025.emnlp-main)

Copied to clipboard

Challenge: Existing model-based channel prediction methods suffer from limited accuracy due to imperfect temporal modeling, while existing AI-based methods suffers from limited generalization due to inadequate training strategies.
Approach: They propose a generative pre-trained language model for channel prediction based on channel correlation and train it based upon transformer decoder architecture.
Outcome: The proposed model can learn various channel characteristics and perform impressive tasks across multiple dimensions.
Enhancing Textbooks with Visuals from the Web for Improved Learning (2023.emnlp-main)

Copied to clipboard

Challenge: Textbooks lack visuals that support student learning, but many lack them . e-textbooks lack such visuals, and many lack these visuals .
Approach: They propose to use vision-language models to automatically enhance textbooks with images from the web.
Outcome: The proposed model improves textbooks with images from the web while allowing for better pedagogical value.
Improving generalization in large langue model by learning prefix subspaces (2023.findings-emnlp)

Copied to clipboard

Challenge: emergence of large language models has significantly transformed the applications of deep learning methods in natural language processing.
Approach: They propose to improve LLMs' generalization by optimizing entire models in parameter space by learning entire simplexes of continous prefixes.
Outcome: The proposed method improves generalization of large language models in the scarce data regime.
Visually Grounded Reasoning across Languages and Cultures (2021.emnlp-main)

Copied to clipboard

Challenge: a new protocol allows for a multilingual hierarchy of concepts and images based on native speakers . the results suggest that the current models are not robust enough to handle multilingual data .
Approach: They propose a protocol to construct an ImageNet-style hierarchy representative of more languages and cultures.
Outcome: The proposed protocol lets the selection of concepts and images be entirely driven by native speakers, rather than scraping them automatically.
This Reads Like That: Deep Learning for Interpretable Natural Language Processing (2023.emnlp-main)

Copied to clipboard

Challenge: In this work, we explore the extension of prototypical networks to natural language processing.
Approach: They propose a weighted similarity measure that enhances the similarity computation by focusing on informative dimensions of pre-trained sentence embeddings.
Outcome: The proposed method improves predictive performance on AG News and RT Polarity datasets and the rationale-based recurrent convolutions.
DetGPT: Detect What You Need via Reasoning (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in the field of computer vision have enabled more effective and sophisticated interactions between humans and machines.
Approach: They propose a reasoning-based object detection paradigm that leverages state-of-the-art multi-modal models and open-vocabulary object detectors to perform reasoning within the context of the user’s instructions and the visual scene.
Outcome: The proposed method enables users to interact with the system using natural language instructions, allowing for a higher level of interactivity.
Beyond Completion: A Foundation Model for General Knowledge Graph Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing foundation models for general knowledge graph reasoning have focused on their structural aspects, with most efforts restricted to in-KG tasks.
Approach: They propose a conditional encoding architecture that bridges the gap between textual and structural modalities, enabling seamless integration.
Outcome: The proposed model outperforms baseline models on 28 datasets and is generalized to out-of-KG tasks.
Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions (2024.emnlp-main)

Copied to clipboard

Challenge: Building socially-intelligent AI agents involves creating agents that can sense, perceive, reason about, learn from, and respond to affect, behavior, and cognition of other agents.
Approach: They propose a set of technical challenges and open questions for researchers to advance Social-AI.
Outcome: The proposed frameworks are based on the social intelligence competencies that evolved over thousands of years in Homo sapiens and are expected to be the foundations for the development of social-intelligent AI agents.
LADDER: Language-Driven Slice Discovery and Error Rectification in Vision Classifiers (2025.findings-acl)

Copied to clipboard

Challenge: Current slice discovery methods in computer vision rely on converting input images into sets of attributes and testing hypotheses about configurations of pre-computed attributes associated with elevated error patterns.
Approach: They propose a method to identify systematic biases in the mistakes of pre-trained vision models by converting input images into sets of attributes and testing hypotheses about configurations of these attributes.
Outcome: The proposed method outperforms existing methods on 3 natural and 3 medical imaging datasets and generates pseudo-labels for each identified bias.
BrainLoc: Brain Signal-Based Object Detection with Multi-modal Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: BrainLoc is a lightweight object detection model guided by fMRI signals.
Approach: They propose a brain-based object detection model guided by fMRI signals . they employ a multi-modal alignment strategy that enhances fmr feature extraction .
Outcome: The proposed model improves fMRI-based object detection accuracy and convenience.
The Sonar Moment: An Audio Geo-Localization Benchmark for Audio-Language Models (2026.findings-acl)

Copied to clipboard

Challenge: AGL1K is the first audio geo-localization benchmark for audio language models (ALMs) it is based on a crowd-sourced platform and is available in 72 countries and territories.
Approach: They propose a benchmark for audio geo-localization that quantifies the informativeness of each recording and a metric that quantizes the information of each audio clip.
Outcome: The proposed benchmarks cover 72 countries and territories and can be used to improve audio geo-localization.
CAVE : Detecting and Explaining Commonsense Anomalies in Visual Environments (2025.emnlp-main)

Copied to clipboard

Challenge: a new benchmark for computer vision fails to capture richness and unpredictability of real-world anomalies . state-of-the-art VLMs struggle with visual anomaly perception and commonsense reasoning . elucidating the nature of anomalies is a fundamental human trait .
Approach: They propose a benchmark for visual anomalies that includes annotations for visual grounding and categorizing anomalies based on their visual manifestations, their complexity, severity, and commonness.
Outcome: The proposed benchmark improves on existing vision models by incorporating visual annotations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations